Papers by Pius Von Däniken
A Measure of the System Dependence of Automated Metrics (2025.acl-short)
Copied to clipboard
| Challenge: | Recent advances in machine translation evaluations are expensive and time-intensive. |
| Approach: | They propose a method to evaluate the correlation between human and metric scores . they argue that it is equally important to ensure that metrics treat all systems fairly and consistently. |
| Outcome: | The proposed method ignores a central requirement of the evaluation process, and ignores the need for a thorough evaluation procedure. |
Favi-Score: A Measure for Favoritism in Automated Preference Ratings for Generative AI Evaluation (2024.acl-long)
Copied to clipboard
| Challenge: | Generative AI systems are becoming ubiquitous for all kinds of modalities . evaluation of generated outputs is increasingly difficult due to cost and complexity of human evaluations. |
| Approach: | They propose to evaluate preference ratings on sign accuracy and favoritism . they propose to use automated metrics to assess generated outputs . |
| Outcome: | The proposed evaluations of preference ratings rely on correlation to human judgments or sign accuracy scores, but this does not tell the whole story. |
Probing the Robustness of Trained Metrics for Conversational Dialogue Systems (2022.acl-short)
Copied to clipboard
| Challenge: | Existing methods for evaluating conversational dialogue systems have been shown to be inefficient and instabile. |
| Approach: | They propose an adversarial method to stress-test trained metrics for evaluation of conversational dialogue systems using Reinforcement Learning. |
| Outcome: | The proposed method outperforms existing methods and can be applied to stress-test trained metrics for conversational dialogue systems. |